Papers with spatial and temporal representations

2 papers
VideoQA-TA: Temporal-Aware Multi-Modal Video Question Answering (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for video question answering align visual or textual features directly with large language models, limiting the deep semantic association between modalities and hindering a comprehensive understanding of interactions within spatial and temporal contexts.
Approach: They propose a temporal-aware framework for multi-modal video question answering that aligns videos and questions at fine-grained levels.
Outcome: The proposed framework improves reasoning ability and accuracy of videoQA by aligning videos and questions at fine-grained levels.
Multi-Channel Spatio-Temporal Transformer for Sign Language Production (2024.lrec-main)

Copied to clipboard

Challenge: Sign language production models ignore structural correlations between channels and use multi-channel spatial attention to capture correlations across channels.
Approach: They propose a novel approach to transform sign language into a unified feature representation using multi-channel spatial attention and temporal attention to learn sequential dependencies for each channel over time.
Outcome: The proposed model outperforms state-of-the-art models on two sign language datasets from diverse cultures.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations